Back

The American Journal of Human Genetics

Elsevier BV

Preprints posted in the last 7 days, ranked by how well they match The American Journal of Human Genetics's content profile, based on 234 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit.

1
A Curated Pharmacogenomic Allele Catalog for Sub-Saharan African Populations

SULAIMAN, M. A.; Oyeyemi, B. F.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361354 medRxiv
Top 0.3%
13.4%
Show abstract

Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.

2
A 515,579-Genome Reference Panel Improves Rare-Variant Imputation Across Multiple Underrepresented Populations

Ivankovic, F.; Ko, A.; Aster, M. M.; Balaconis, M. K.; Banks, E.; Bemis, M.; Cibulskis, K. R.; Degatano, K.; Gauthier, L. D.; Grant, G.; Hatcher, A.; Kachulis, C.; Karczewski, K. J.; Labrecque, S. M.; Lawson, J.; Liao, C.; Magner, R.; Munshi, R.; Schatz, M. C.; Schultz, P. M.; Shah, S. P.; Sheets, E. A.; Tibbetts, K.; Vernest, K. A.; Ye, R.; Gabriel, S.; Lennon, N. J.; Neale, B. M.; Browning, B. L.; Lichtenstein, L. T.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361247 medRxiv
Top 0.5%
9.5%
Show abstract

Genotype imputation remains essential for large-scale human genetics studies, but its performance is limited by the size and ancestral diversity of available reference panels, reducing accuracy for rare variants and underrepresented populations. Here, we present a cloud-based imputation service built on a multi-ancestry reference panel derived from 515,579 jointly phased genomes from the All of Us (N=414,830) and National Human Genome Research Institute's Analysis, Visualization, and Informatics Lab-space (AnVIL, N=100,749) datasets. The All of Us + AnVIL reference panel is highly diverse and includes 261,163 participants most genetically similar to non-European reference populations, spanning 665,398,839 high-quality autosomal sites, representing a nearly 50% increase over TOPMed, the previous largest imputation service. Across multiple ancestry groups, the panel enables accurate imputation (empirical R2 0.8) for variants with allele frequencies as low as 0.2%, extending reliable imputation into the rare-variant frequency spectrum, including allele frequencies down to 0.002% and 0.006% for samples with European ancestry and African ancestry in the United States, respectively. Compared with TOPMed, the panel improves imputation accuracy across all ancestry groups except Africans, and recovers additional trait-associated variants not represented in existing reference panels. To facilitate broad community access while preserving participant privacy, we deploy the panel through a secure cloud-based imputation platform using privacy-preserving recombined haplotypes. This resource establishes a new foundation for genome-wide association studies (GWAS) and fine-mapping, especially in previously underrepresented populations.

3
A novel framework leveraging non-causal associations reveals shared pathways linking inflammation and cancer risk

Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.30.26361622 medRxiv
Top 0.6%
8.9%
Show abstract

Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.

4
Constitutive PDGFRb activation drives connective tissue overgrowth through STAT5-IGF1 signaling

Kwon, H. R.; Rackley, A.; Olson, L. E.

2026-08-29 genetics 10.64898/2026.08.27.747555 medRxiv
Top 0.6%
7.8%
Show abstract

Autosomal dominant gain-of-function mutations in platelet-derived growth factor receptor beta (PDGFRb) cause overgrowth of the skeleton and other connective tissue in Kosaki overgrowth syndrome. However, the target cell type and signaling pathways underlying PDGFRb-driven overgrowth are unknown. Normal postnatal growth is controlled by pituitary-secreted growth hormone (GH), which activates the STAT5 transcriptional factor to upregulate insulin-like growth factor 1 (IGF1). To investigate the role of the GH-STAT5-IGF1 pathway in PDGFRb-related overgrowth, we generated mice with a PDGFRb gain-of-function mutation in skeletal and fibroblast lineages, which resulted in STAT5 activation and gigantism. Conditional deletion of Stat5ab in connective tissue lineages rescued skeletal overgrowth and keloid-like fibrosis in the skin. Conditional deletion of GH receptor (Ghr) did not rescue overgrowth, indicating the physiological activator of STAT5 is not required for overgrowth. However, deletion of Igf1, the STAT5 target gene, and its receptor, Igf1r, in connective tissue, rescued the overgrowth phenotype. These findings demonstrate a GHR-independent STAT5-IGF1 signaling pathway in mutant connective tissue cells, which mediates PDGFRb-driven overgrowth in mice and potentially in humans with similar PDGFRB mutations.

5
ICONIC: An R Package for Integrating Instrumental Variable- and Negative-Control-Informed Causal Discovery and Diagnostics in Multiomic Studies

Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26361466 medRxiv
Top 0.7%
7.2%
Show abstract

Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.

6
Efficient genome-wide mapping of reproducible, context-dependent eQTLs at single-cell resolution

Alquicira-Hernandez, J.; Dorans, E.; Tomofuji, Y.; Nathan, A.; Raychaudhuri, S.

2026-08-29 genetics 10.64898/2026.08.25.747138 medRxiv
Top 0.7%
7.2%
Show abstract

Single-cell technologies enable linking disease-risk variants to gene regulatory effects in specific cell-state contexts. However, most so called "single-cell eQTL" studies use a "pseudobulking" strategy to identify expression Quantitative Trait Loci (eQTLs), obscuring subtle dynamic regulatory effects of disease alleles. Here, we propose Dynema (Dynamic eQTL mapping in single cells) for fast and accurate genome-wide mapping of context-dependent and independent eQTL effects at true single-cell resolution. To identify eQTLs, Dynema uses a Poisson model with cluster robust variance estimators (CRVEs) to account for correlation of single-cell profiles from the same individual. In contrast to other common methods, Dynema achieves statistical calibration and scales to genome-wide analysis in large single-cell datasets in realistic timeframes. We applied Dynema to two independent T cell datasets and identified reproducible cell-state-dependent eQTL effects. Some cell-state-dependent eQTLs are missed by pseudobulking approaches, and many others are conditionally independent from lead eQTL effects. We show that TSPAN32 and other autoimmune loci colocalize with cell-state-dependent eQTLs. Mapping context-dependent eQTLs at single-cell resolution enables the definition of the molecular effects of complex disease alleles.

7
Loss of RUBCN causes autophagy overdrive in a neurodevelopmental disorder with age-dependent neurodegeneration

Efthymiou, S.; Tabata, K.; Dafsari, H. S.; Schober, E.; Latza, C.; Isaoglu, M.; Abuelrub, A.; Rad, A.; Firoozfar, Z.; Turchetti, V.; Lin, R. Q.; Maroofian, R.; Wiethoff, S.; Afzal, E.; Zafar, F.; Rana, N.; McRae, A. M.; Kaiyrzhanov, R.; Guliyeva, U.; Gulieva, S.; Melikishvili, G.; Lespinasse, J.; Vitobello, A.; Denomme-Pichon, A.-S.; Wentzensen, I. M.; Mefford, H. C.; Briere, L. C.; A Walker, M.; A High, F.; Sweetser, D. A.; Kendall, M.; Franchi, M.; Brown, M.; Latner, D.; Joset, P.; Ivanovski, I.; Alfadhel, M.; Alluhaydan, I.; Frederiksen, A. S.; Arriens, V.; Hanker, B.; Mankad, K.; Guerin, J

2026-09-01 genetic and genomic medicine 10.64898/2026.08.27.26360298 medRxiv
Top 1%
3.5%
Show abstract

Pathogenic variants in RUBCN, encoding the Run domain Beclin-1 interacting and cysteine-rich domain-containing protein (Rubicon) have been implicated in autosomal recessive spinocerebellar ataxia 15 (SCAR15). However, the molecular mechanisms underlying disease pathogenesis remain poorly understood. Here, we report 18 individuals from 15 unrelated families harbouring biallelic RUBCN variants, who present with an aggressive neurodevelopmental disorder variably characterized by seizures, developmental delay, intellectual disability and movement abnormalities that cause regression, progressive brain atrophy and neurodegenerative features. Through functional characterization, we demonstrate that a subset of disease-associated putative truncating variants disrupt autophagy regulation. In Caenorhabditis elegans models, loss-of-function RUBCN variants result in an increased autophagic flux and impaired neuronal function, recapitulating key features in humans. Correspondingly, cellular assays reveal that nonsense and frameshift RUBCN variants lead to defective autophagy inhibition, underscoring a crucial role for RUBCN as a key negative autophagy regulator. Molecular dynamics simulations rank the eleven missense variants by structural effect, with p.Arg813Trp alone altering the target protein at both the local and the regional level and lying within the RAB7A-binding module that the truncating alleles remove altogether. Our findings establish and expand the RUBCN-related disorders as a clinically and molecularly distinct subset of autophagy-related diseases. By delineating both the genetic landscape and cellular consequences of Rubicon dysfunction, this study enhances our understanding of autophagy-related neurodevelopmental disorders and provides a foundation for future therapeutic investigations.

8
A comprehensive atlas of somatic mutation rates and mutational signatures in normal human cells

Pham, M. H.; Harvey, L. M. R.; Oliver, T. R. W.; Dunstone, E.; Lawson, A. R. J.; Nicola, P. A.; Sanghvi, R.; Hooks, Y.; Mitchell, E.; Jarman, G. L.; Wang, Y.; Abascal, F.; Jung, H.; Neville, M. D. C.; Ishida, Y.; Fowler, J. C.; Le, A. P.; Moody, S.; Marshall, H.; Brzozowska, N.; Ding, C.; Pac, C. A.; Machado, H. E.; O'Neill, L.; Latimer, C.; Humphreys, L.; Saeb-Parsy, K.; Mahbubani, K. T. A.; Baxter, J.; Rassl, D. M.; Vicario, R.; Geissmann, F.; Kabashima, K.; Bleys, R. L. A. W.; Moore, L.; Heer, R.; Coorens, T. H. H.; Behjati, S.; Hoare, M.; Campbell, P. J.; Jones, P. H.; Martincorena, I.; Ra

2026-08-29 genomics 10.64898/2026.08.28.747772 medRxiv
Top 1%
3.3%
Show abstract

Over the course of a lifetime, somatic mutations accrue in normal human cells, causing variation in cell phenotype and engendering somatic evolution with outcomes ranging from the adaptive immune system to cancer. To inform understanding of somatic evolution in the human body we report the mutation rates and mutational signatures of 53 normal cell types. Most show evidence of linear mutation accumulation over time with single base substitution mutation rates ranging from ~3.5/year/diploid genome in spermatogonia and sperm, to ~20/year in postmitotic neurons, ~50/year in mitotically active colorectal epithelial cells, ~60/year in kidney proximal tubule cells and hepatocytes, 100s/year in sun-exposed skin epidermal cells and 10-50/year in the remainder. Certain cell types, including skin epidermis, cardiac myocytes, bladder urothelium, kidney proximal tubule cells, and hepatocytes, show substantial variability in mutation burdens around the linear age trend, indicating the influence of additional factors which differ between individuals and modulate mutation accumulation, including exogenous mutagen exposures. At least 18 single-base substitution and nine small insertion and deletion mutational signatures are present, some in all cell types, some in a subset and others in a single cell type. Known exogenous mutagen exposures and endogenous mutational processes account for some mutational signatures, but the origins and mechanisms underlying many are uncertain. This comprehensive survey of mutagenesis provides a foundation for understanding somatic evolution of human cell populations in health and disease.

9
Deep phenotyping and multi-omics analyses reveal systems-wide metabolic dysregulation in a refined trisomy mouse model of Down syndrome

Saqib, M.; Chen, F.; Mistri, D. K.; Tan, L.; Wright, N.; Sarver, D. C.; Anders, R.; Aja, S.; Wong, G. W.

2026-08-29 physiology 10.64898/2026.08.26.747201 medRxiv
Top 2%
2.5%
Show abstract

Trisomy 21 or Down syndrome (DS) affects multi-organ systems across the lifespan. The presence of an extra chromosome, along with genome dosage imbalance due to triplicated genes, contributes to the DS phenotypes. Of the DS mouse models, few are aneuploid with a freely segregating extra chromosome. We previously showed that the aneuploid Ts65Dn mice exhibit metabolic deficits consistent with the metabolic profile of DS. However, the genotype-phenotype relationships in Ts65Dn mice are complicated by the presence of triplicated genes unrelated to human chromosome 21 (Hsa21). To address this issue, we leveraged a refined model, Ts66Yah, where the extra triplicated genes in Ts65Dn have been removed. Deep phenotyping and multi-omics analyses showed that Ts66Yah mice develop pronounced and widespread metabolic disturbances. Despite sexual dimorphism in weight gain, body temperature, lipid and lipoprotein profiles, hepatic injury and adipose fibrosis, both male and female Ts66Yah mice share a common phenotype of pronounced glucose intolerance and insulin resistance, reduced mitochondrial respiratory capacity in visceral fat, altered serum inflammatory cytokine profile, and dysregulated serum and liver metabolomes. Pan-tissue transcriptomes also reveal signatures of immune activation, disrupted metabolic processes and cellular respiration, altered cytokine signaling, enhanced oxidative stress, and extracellular matrix remodeling. These combined changes across tissues disrupt metabolic homeostasis more severely in Ts66Yah than in Ts65Dn mice. Several phenotypes, including glucose intolerance, insulin resistance, tissue fibrosis, and oxidative stress were further exacerbated by an obesogenic diet. This foundational data establishes Ts66Yah as a valuable reference model for the mechanistic and comparative study of metabolic dysfunction in DS.

10
Young people with obesity and rare disease - genotypes, phenotypes and healthcare use

Chia, C.; Baker, K.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361359 medRxiv
Top 2%
2.4%
Show abstract

Obesity is a significant public health concern. Early-onset obesity in the context of rare disease can reflect genetically-mediated pathology or elevated susceptibility through indirect mechanisms. Mapping the diverse characteristics and needs of young people with obesity in the rare disease population is a first step toward mechanistic and translational research. We carried out a retrospective comparative analysis of demographic, genotypic, phenotypic and health service utilisation data for young people with obesity (cases: n=500) and without obesity (controls: n=11,444) from the UK 100,000 Genomes Project rare disease cohort. Cases and controls were recruited prior to genomic diagnosis, across clinical disorder categories. We observed significant association between socioeconomic deprivation and obesity risk. Young people with obesity had significantly higher utilisations of acute care and mental health services, indicating an overall higher health burden. A curated panel of 519 candidate obesity-associated genes demonstrated aggregate association with obesity, although no single gene reached significance. Phenotypic comparison between cases and controls highlighted increased multi-organ and neurological system involvement, highlighting the overlap between neurodevelopmental and obesity risks. Within the case group, we conducted cluster analysis to identify early-onset obesity groups with different phenotypic profiles, potentially arising from different causal pathways - this identified six obesity subgroups of interest, with differing involvement of neurodevelopmental and other systems. Our study confirms that obesity co-occurs with a wide range of factors within the rare disease population, and is associated with significant physical and mental health needs, requiring holistic lifelong care.

11
Genomic Architecture of Migraine: A Multi ancestry GWAS Meta analysis of 2.5 Million Participants

Overstreet, C.; Galimberti, M.; Harsan, K. T.; Beck, S. E.; Hirsch, J.; Sariya, S.; Ferolito, B. R.; Zhou, Y.; Zhang, Y.; Weinheimer, E. I.; Lacobelle, A.; Nunez, Y.; The VA Million Veteran Program, ; Kranzler, H. R.; Gaziano, J. M.; Stein, M.; Gottschalk, C.; Choi, K. W.; Pereira, A. W.; Deak, J. D.; Pathak, G. A.; Levey, D. F.; Gelernter, J.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.28.26361638 medRxiv
Top 2%
2.4%
Show abstract

Migraine is a leading cause of disability, yet preventive treatment remains largely empirical despite the availability of several mechanistically distinct therapies. Genetic data can clarify mechanisms and therapeutic hypotheses when association signals are integrated with molecular and clinical data. We meta-analyzed migraine GWAS data from 12 European ancestry cohorts (206,893 cases and 2,093,175 controls) and four African ancestry cohorts (22,115 cases and 178,626 controls). We identified 311 lead variants in European-ancestry analyses and 316 lead variants in trans-ancestry analysis. Fine-mapping and transcriptome-wide analyses prioritized variants and genes implicated in sensory neuronal signaling, vascular tone, and immune regulation, with convergent evidence at several established loci including TRPM8 and PHACTR1. Drug-repurposing analyses identified therapeutic targets and compounds, including established migraine treatments and candidates requiring experimental validation. Genetic correlations, Mendelian randomization, and a phenome-wide scan linked migraine liability to psychiatric, pain, and gastrointestinal phenotypes. Together, these findings expand the known genetic architecture of migraine across ancestries and provide a genetics-led map connecting association signals with biological pathways, multimorbidity and candidate therapeutic mechanisms, providing a foundation for future functional and translational studies.

12
A Randomized Non-Inferiority Trial of an eHealth Delivery Alternative for Cancer Genetic Testing for Hereditary Cancer (eREACH2)

Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26361920 medRxiv
Top 2%
2.2%
Show abstract

Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.

13
Rare and Common Germline and Somatic Variants Shape Immune Cytopenia Risk and Enable Risk Stratification

Faria, S. D. S.; Bineau, J.; Moisan, R.; Legault, M.-A.; Lecluze, E.; Pincez, T.

2026-08-31 hematology 10.64898/2026.08.26.26361484 medRxiv
Top 2%
1.9%
Show abstract

The genetic risk factors of immune cytopenias are unclear. Immune cytopenias have been reported in various genetic contexts: 1) inherited error of immunity genes, mainly due to rare germline variants, 2) systemic lupus erythematosus, associated with common germline variants, 3) hematological malignancies, and 4) clonal hematopoiesis, the latter two due to somatic variants. However, the respective contribution and interaction of these variants remain to be investigated. Here, we used two large biobanks with whole genome sequencing data to systematically investigate the genetic contribution to immune cytopenia. We found that the four types of genetic variants independently contribute to immune cytopenia risk. We notably found that carriers of variants in some autosomal recessive genes of inherited error of immunity had an increased risk of immune cytopenia. Additionally, common variant-mediated risk of systemic lupus erythematosus also increased the risk of immune cytopenia. Overall, a third to a half of patients with immune cytopenia carried at least one of the four genetic risk variants investigated. Combining the four variants allowed stratifying the risk of immune cytopenia in both general and high-risk population. In general population, the 10-year incidence of immune cytopenia in the lowest and highest risk groups was 0.08% and 1.5%, respectively. In sum, this work identified that different genetic risk factors can lead to immune cytopenia. A large proportion of individuals with immune cytopenia carried an underlying genetic risk factor. Finally, combining these genetic risk factors enabled risk stratification.

14
Molecular Underpinnings of Retinal Traits 1 Shared with Major Psychiatric Disorders

Jaholkowski, P.; Parker, N.; Sveen, I. O.; Wistrom, E. D.; Fominykh, V.; Szabo, A.; Parekh, P.; Frei, O.; Smeland, O. B.; O'Connell, K. S.; Djurovic, S.; Dale, A. M.; Shadrin, A. A.; Andreassen, O. A.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.31.26361809 medRxiv
Top 2%
1.5%
Show abstract

Recent large-scale studies have enabled new knowledge about genetic underpinnings of morphological and electrophysiological alterations of the retina. Variation in retinal traits, often of neurodevelopmental origin, have been linked to major psychiatric disorders (MPDs). Here, we investigate the genetic overlap between MPDs and key retinal traits to identify underlying molecular mechanisms. We obtained genome-wide associations studies data for bipolar disorder (BD), major depression (MD), schizophrenia (SCZ), and the retinal traits retinal nerve fibre layer thickness (RNFL), ganglion cell inner plexiform layer thickness (GCIPL), and vertical cup-disc ratio (VCDR). We estimated the number of trait-influencing variants shared between traits with MiXeR and identified shared genetic loci with condFDR. Subsequently, we examined the biological pathways of the genes mapped to shared loci. This revealed that GCIPL shared the most genetic variants with MPDs (~60%), followed by RNFL (~40%), and VCDR (~20%). The genetic variants shared between retinal traits and MPDs showed disorder-specific patterns with more pronounced overlaps of SCZ and BD with RNFL, and MD negatively correlated with GCIPL. Gene-pathway analysis highlighted the importance of GABAergic neurotransmission and a two-stage neurodevelopmental process in SCZ, whereas the role of mitochondria and a weaker developmental component were observed in BD. The results also implicated synaptic functioning and gene-expression processes in MD. Furthermore, polygenic analysis suggested that the genetic architecture of retinal traits can distinguish between MPDs. Our findings indicate shared genetic underpinnings between retinal traits and SCZ, BD, and MD, implicating altered neurodevelopment and neurotransmission underlying the retinal link to major psychiatric disorders.

15
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference

Alve, S. R.; Rahman, S.; Meem, S. M. A. C.

2026-09-02 dentistry and oral medicine 10.64898/2026.09.01.26361874 medRxiv
Top 2%
1.5%
Show abstract

A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.

16
A germline KDM3C polymorphism impairs DNA repair and sensitizes to chemoradiotherapy

Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26360896 medRxiv
Top 2%
1.4%
Show abstract

Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.

17
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 2%
1.3%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

18
Clinical deep sequencing to diagnose pathogenic mosaic variants in malformations of cortical development and epilepsy

Stone, K.; Prinzing, G.; Lai, A.; Smith, L.; Sheidley, B. R.; Corliss, M. M.; Bowling, K.; Cao, Y.; Wiltrout, K.; Stone, S. S. D.; Lidov, H.; Yang, E.; Poduri, A.; D'Gama, A. M.

2026-09-03 neurology 10.64898/2026.09.01.26361943 medRxiv
Top 3%
1.0%
Show abstract

Background and Objectives: Deep sequencing of brain tissue in the research setting has established that mosaic variants are a major cause of malformations of cortical development (MCDs) and epilepsy. However, genetic testing in the clinical setting primarily detects germline variants using clinically accessible samples. We aimed to determine the diagnostic yield and clinical utility of deep sequencing in the clinical setting to identify pathogenic mosaic variants for this population. Methods: We performed a retrospective cohort analysis of individuals at Boston Children's Hospital with MCDs with or without epilepsy who received clinical deep sequencing between September 2017 and February 2026. Demographic, clinical, and genetic testing data were abstracted from the medical record. For individuals without systemic features, we classified brain tissue as an affected tissue sample. For individuals with systemic features, we classified brain or relevant non-brain tissue as affected. The primary outcome was the diagnostic yield of clinical deep sequencing performed using affected vs unaffected tissue samples. The secondary outcome was the clinical utility of genetic diagnoses. Results: Our cohort included 37 individuals (19/37 (51%) female, 18/37 (49%) male) with MCDs, of whom 35/37 (95%) had epilepsy (25 with brain tissue samples available from epilepsy surgery) and 8/37 (22%) had systemic features. Most (35/37 (95%)) had dysplasia phenotypes on MRI and 12/27 (44%) with pathology available had Focal Cortical Dysplasia Type I or II. The diagnostic yield was 53% (17/32; 16 mosaic and 1 germline variant) when clinical deep sequencing was performed using an affected tissue sample vs 0% (0/6) using an unaffected tissue sample (p=0.016). Of the diagnosed cases, 13/17 (76%) had testing performed on brain tissue (1 with systemic features) and 4/17 (24%) on non-brain tissue (3 buccal and 1 duodenal tissue, all with systemic features). All but one diagnosis involved the mTOR pathway. All diagnoses had clinical utility. Discussion: Clinical deep sequencing, when performed using an affected tissue sample, has high diagnostic yield and clinical utility for individuals with MCDs, especially dysplasia phenotypes, and epilepsy. Our findings support implementation of clinical deep sequencing for this population, especially as the genetic diagnoses have implications for emerging precision therapies.

19
Selection and surveillance of 5S ribosomal RNA genes in human populations

Sengl, L.; Bagaric, I.; Conil, C.; Seeleuthner, Y.; Mueller, M.; Klughammer, J.; Mages, S.; Cobat, A.; Bohlen, J.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.27.26361558 medRxiv
Top 3%
0.9%
Show abstract

The 5S ribosomal RNA gene is present in the human genome not once but in ~80 copies, arranged head to tail in a single array of ribosomal DNA on chromosome 1 -one of the most repetitive and least explored regions of the genome. Its product is one of the four RNAs in every ribosome and, when ribosome assembly fails, it activates the tumour suppressor p53. Whether these copies vary in sequence between people, and whether such variation has physiological or pathological consequences, is unknown. Using telomere-to-telomere genome assemblies, whole-genome sequences from ~490 000 UK Biobank participants, and ~940 GTEx transcriptomes, we find that every person carries copies bearing substitutions or indels, and that ~10% of people express such variant 5S rRNA. Mutating every position of the gene in vitro, we find that variants blocking incorporation into the ribosome map to the uL5/uL18 interface and activate p53. Remarkably, these same variants are depleted from human populations: selection has acted on the step that p53 monitors. Ribosomal DNA is thus a functional source of human genetic variation, long invisible to genome-wide analysis and shaped by the p53 pathway it controls.

20
Molecular landscape and risk stratification in acute myeloid leukemia - insights from the real-world REFORM-AML cohort

Kristensen, D. T.; Broendum, R. F.; Knudsen, M.; Grubach, L.; Marcher, C.; Preiss, B.; Bibi, M. L.; Hoegdall, E.; Poulsen, T.; Skov, V.; Oerskov, A. D.; Groenbaek, K.; Hansen, J. W.; Schoellkopf, C.; Cowland, J.; Andersen, M. K.; Severinsen, M. T.; Vejgaard, C.; Larsen, O. H.; Vang, S.; Boegsted, M.; Roug, A. S.

2026-08-31 hematology 10.64898/2026.08.27.26361552 medRxiv
Top 3%
0.6%
Show abstract

Large genomically annotated acute myeloid leukaemia (AML) datasets exist, but population-based contemporary cohorts remain scarce. Here we report clinicopathological, genomic, and outcome data from Danish AML patients. 2,512 AML patients were identified between 2015-2022, of whom 33.8% had available NGS data (NGS+). In patients [≤]70 years, baseline characteristics and outcomes were comparable between NGS+ and NGS- groups. In patients >70 years, more NGS+ patients received intensive treatment, but survival was similar among intensively treated patients. The distribution of mutations varied significantly by age and sex, with older age and male sex exhibiting higher frequencies of adverse-risk gene mutations. In intensively treated NGS+ patients, ELN2017 stratified 5-year OS: 58.4% (favorable), 43.4% (intermediate), and 28.2% (adverse), with hazard ratios (HRs) of 0.63 (favorable) and 1.45 (adverse) relative to intermediate. ELN2022 yielded corresponding OS rates of 56.9%, 51.8%, and 29.7%, with HRs of 0.78 and 1.86. The two models had comparable predictive performance for OS in a time-dependent model. In conclusion, outcomes of intensively treated AML patients were comparable irrespective of NGS status, underscoring the representativeness of the REFORM-AML database for the Danish AML population. Age and male sex correlated with adverse-risk mutations, and both ELN2017 and ELN2022 robustly predicted survival.